Skip to content

feat: gate managed Turn host starts by owner cadence - #4929

Merged
huangruiteng merged 8 commits into
mainfrom
codex/turn-atomic-cadence-admission
Sep 24, 2026
Merged

huangruiteng merged 8 commits into
mainfrom
codex/turn-atomic-cadence-admission

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

An owner can set a 24-hour automatic execution floor, yet a managed turn run-once host failure or invalid result could start a new model immediately on retry. This change reserves a start atomically in the existing quota policy store before each new managed host attempt. An early retry returns the next eligible time without invoking the host, writing back a result, or spending a quota slot; cached settlement still replays without another reservation.

The store upgrades M1 policy files in place to v2 and retains start records across restart. Goal floors apply per agent, and an automation-specific floor requires a stable --automation-id. An explicit --manual-interval-bypass-reason bypasses only the interval and records the manual start. The CLI renders the interval wait and next eligible time. The English and Chinese RFCs now distinguish this managed Turn candidate from the still-unqualified Codex App timer-to-hook path and other launchers.

Validation: 140 Python Turn/cadence tests; six TypeScript cadence tests covering exact due time, concurrency, restart, scope isolation, manual bypass and M1 migration; focused final CLI/retry tests; TypeScript typecheck, Ruff, public-boundary scan, and exact-head change-quality receipt cqr_a7e80b9285d5ac8e1ad9 passed. The risk-based premerge gate passed all 19 checks with no manual holds. The scoped mypy command has existing errors on main (87) and this branch (84), with no new normalized errors. No live model call or existing automation change was made.

This PR changes the managed CLI entry point and quota/Turn readback. The Codex App schedule recommendation remains the M1 behavior; App pre-model hook coverage and the packaged settings editor are separate acceptance work. No frontend code changes are needed for this managed CLI slice.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Request changes conclusion (author-owned PR; GitHub blocks formal self-review)

动机

这项改动试图把 owner 设定的最小自动启动间隔,从仅影响调度建议推进为 managed turn run-once 的执行前硬门槛。旧路径在失败后立即重试时仍可能启动 host。目标有价值,但本 head 还不能证明失败/崩溃后的 Turn 可以按间隔安全恢复。

改动思路

TypeScript cadence store 在同一文件锁下处理规则和启动预留;Python Turn executor 在 host 与账本写入前请求 admission。未配置 policy 时走 default-off 直通,不创建 store。这个 owner 划分和 per-agent/automation scope 与现有架构相符;问题发生在 TS 预留成功与 Python 首次持久化 Turn 尝试之间的跨边界交接。

具体改动

关键代码讲解

  • admitAutomationStart (loopx/control_plane/quota/automation_cadence.ts) 在锁内保存 request_id 和启动时间,并把相同 request_id 一律判为 duplicate_or_stale_trigger,即使下一次调用已过最小间隔。
  • admit_managed_start (loopx/cli_commands/turn.py) 用 turn_key:attempt 生成 request ID;对同一个尚无 journal/attempt 记录的 Turn,重启会再生成 turn_key:1。
  • run_loopx_turn_once (loopx/control_plane/turn_driver/executor.py) 先调用 admit_start,之后才创建或更新 journal;host_attempt_count 又要到 host stage 才递增。被 admission 拒绝时直接返回 interval_wait,host、writeback 和 quota 均未执行。

本轮在 exact head 998c79535711361594a8f141e4809d3111b8f63b 上跑过 138 项 Python Turn 测试、6 项 TypeScript cadence 测试、TypeScript typecheck、Ruff 和 diff hygiene,均通过;风险型 premerge 另有 19/19 项通过,公开边界扫描无命中、无 skip/人工 hold。现有 TS 测试也明确断言相同 request ID 过了一个完整间隔仍被拒绝;这些绿灯不能覆盖下述跨边界崩溃窗口。未查询远端 CI,按该 Goal 的 wait_for_ci=false 执行。

对主干的风险

[P1] 已预留但未记入 Turn journal 时,同一个 Turn 无法恢复。 触发序列:TS 成功原子写入 turn_key:1;进程在 executor.py 写入新 journal(或递增 host_attempt_count)前退出,或者 effect RPC 已提交但响应丢失。重启同一 Turn 时仍请求 turn_key:1,TS 在检查 interval 之前以 duplicate_or_stale_trigger 拒绝;即使等满间隔或提供手动 bypass reason,重复 ID 检查仍先返回拒绝。executor 每次都返回 interval_wait,没有 host/账本动作,也没有可供正常恢复的已持久化 attempt。创建另一个 Turn key 可以绕开,但等于放弃原 Turn,并非本 PR 声称的重放/恢复语义。

最小修复是让 admission 与 journal 的交接具备可恢复的明确状态/身份:区分“已预留但尚未发起 host”和“host 可能已发起”,在安全前提下允许前者按原 Turn 恢复,后者保留 fail-closed;或提供显式、可审计的人工恢复结果,而不是无限 interval_wait。请加入真实 TS store + Python Turn 入口的故障注入回归:预留已落盘,分别在首次 journal 写入前和 host_attempt_count 写入前中断,重启并跨过 floor 后验证明确恢复/停止结果、不会重复启动 host、也不会额外 spend。

我的整体评价

结论:REQUEST_CHANGES。正常路径、默认关闭路径、范围及持久化迁移设计均有聚焦验证,未来向的 TypeScript 单一 cadence owner 也合理;但跨边界 crash consistency 是这项硬门槛最关键的负例,不能以单层的原子写入与绿灯测试替代。请修复上述窗口并在新 head 重跑跨边界回归和风险型 premerge 检查;本评论不授权合并。

English verdict: REQUEST_CHANGES - A committed cadence reservation can precede the first durable Turn attempt, leaving the same Turn key permanently rejected as a duplicate after a crash or lost RPC response. Add recoverable handoff semantics and a real TS-store/Python-entrypoint fault-injection regression before re-review.

A committed reservation could precede the first durable Turn journal attempt, so
the same Turn asked again with the same request identity and was rejected as a
duplicate forever, even past the owner floor and with a manual reason.

Managed starts are now two-phase in the same cadence store: admission writes a
reserved start, and the Turn executor confirms it only after the host attempt is
durable in the journal. The same identity resumes a reserved start once the floor
is reached, while a confirmed start stays fail-closed. A record without the phase
field is read as an attempted start, so older or hand-edited stores fail closed.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Keep both new effect imports and routes (this branch's cadence confirm route plus main's shadow-drain and canonical-snapshot page routes), and take main's already-repaired bilingual mirror declaration while keeping this branch's updated implementation baseline.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

上一轮评审在 998c79535711361594a8f141e4809d3111b8f63b 上请求修改,理由是:TS 已原子写入启动预留后,进程可能在 Turn journal 记录 host 尝试之前退出(或 effect 响应丢失),于是同一个 Turn 重启后仍以同一 request ID 申请准入,被 duplicate_or_stale_trigger 永久拒绝——即使等满间隔或给出手动理由也无法恢复,executor 每次都只返回 interval_wait。

本 head e5f70a230ba5c75c484f11f168bf88e4fcb33d6a 修复了该窗口,并合入最新 main。修复方式是让预留与 Turn journal 的交接具备显式、可恢复、可审计的状态,而不是依赖"单层原子写入即安全"的假设。

改动思路

规则所有权不变:TypeScript cadence store 仍是策略与启动记录的唯一 owner,准入在同一 policy 文件锁内完成;Python Turn executor 在任何 host 或 journal 尝试变更之前请求准入,被拒绝时只返回 interval_wait(含下次可执行时间),不调用 host、不写回、不花 quota。

新增的关键区分是"已预留但 host 尚未被持久记录"与"host 尝试已落盘"。启动记录现在带阶段字段:准入写入 reserved,Turn executor 只有在把 host 尝试写入 journal(host_attempt_count 在调用 host 之前先落盘)之后才调用新的 confirm_start 把它置为 started。因此:

  • 同一 request 身份遇到 reserved 记录时,在满足 owner 下限后按原 Turn 恢复(间隔锚点保持首次预留时间,不会被推后或前移);
  • 遇到 started 记录时保持 fail-closed,重复身份与手动理由都不能绕过;
  • 缺少阶段字段的旧记录按"已尝试启动"读取,旧版或手工改写的 store 因此 fail-closed 而不是被当作可恢复预留。

未配置策略时两阶段都不创建 store 或锁文件,默认关闭路径与之前一致。

具体改动

12 文件 +787/−51。TypeScript:automation_cadence.ts 的启动记录增加 reserved/started 阶段并新增 confirmAutomationStart;effect_runtime_handlers.ts 注册 quota.automation_cadence.confirm_start。Python:turn.py 把原先的内联准入闭包提取为可复用的 managed_cadence_start 工厂(准入 + 确认),executor.py 在 host 尝试落盘后、调用 host 前执行确认。测试:TS cadence 用例覆盖恢复/确认/升级/fail-closed;tests/test_loopx_turn_executor.py 新增两个真实 TS store + 真实 Turn 入口的故障注入回归。文档:两份 RFC 记录两阶段交接与恢复语义。本 head 同时合入最新 main(保留双方新增的 effect import/路由,采用 main 已修好的镜像声明,并保留本分支更新的实现基线)。

关键代码讲解

  • admitAutomationStart(loopx/control_plane/quota/automation_cadence.ts):在同一 policy 锁内判定重复/过期触发、owner 下限与手动绕过;同身份的 reserved 记录在满足下限后返回 resumed_unstarted_reservation,started 记录则维持 duplicate_or_stale_trigger。
  • confirmAutomationStart(同文件):幂等确认;无对应预留时返回 reservation_missing 且不创建文件;state 字段缺失的记录在读入时按 started 处理。
  • ManagedCadenceStart / managed_cadence_start(loopx/cli_commands/turn.py):生产 CLI 与回归测试共用同一准入/确认实现,避免测试自建一套旁路语义。
  • _host_result_stage 的确认点(loopx/control_plane/turn_driver/executor.py):host 尝试计数落盘后、host 调用前执行 confirm_start;确认失败即停止,不会在状态未明时启动 host。
  • 故障注入回归(tests/test_loopx_turn_executor.py):分别在"预留后、任何 journal 写入前"与"admission 记录已写、host 尝试计数写入前"中断进程,重启并跨过下限后断言恢复提交、host 仅一次、spend 仅一次、记录最终为 started。

对主干的风险

上一轮点名的 P1 场景现在有真实 TS store + 真实 Python 入口的负例覆盖,两个中断点都验证为可恢复且不重复启动 host、不额外 spend。恢复路径的安全前提是同一 Turn lane 由 journal 锁串行执行、且尝试计数先于 host 调用落盘;这条前提已作为后续风险记录在案,删除该锁或把 host 调用提前都会重新打开重复启动窗口。

仍保持 fail-closed 的部分:已确认启动对同一身份拒绝重复;缺阶段字段的 store 视为已启动;确认失败会以显式失败结束而不是静默继续。范围边界未扩大:App 定时器到 hook、非 Turn launcher、打包设置界面与真实宿主推广仍未验收,本 PR 不晋升任何 provider 或自动化。

验证矩阵(本 head):Python Turn executor/driver/automation cadence/turn envelope/writer fence 共 217 用例通过(含两个崩溃恢复回归);TS cadence 用例 7 项通过;npm run typecheck:control-plane 通过;TS 全量 2936/2961 通过,唯一失败 tests/control_plane_ts/sqlite_capacity.test.ts 在未包含本 PR 的 origin/main 上同样失败(既有环境问题);census 与 goal-instance inventory 9 项通过;docs-governance-smoke.py、配置范围内 Ruff 与 python -m mypy(23 模块)通过。按本 Goal 的 wait_for_ci=false 与维护者要求,本轮未抓取或等待远端 CI;合并前仍对未变 head 运行 merge-readiness。

语义与 CI 对齐

cadence store 由 v1 原地升级为 v2 并新增阶段字段,路径不变,旧版程序拒绝 v2 而不是静默丢弃启动记录;新增的 quota.automation_cadence.confirm_start 属内部 coordination 契约,不改公共 CLI 参数与默认行为。文档已在两份 RFC 中同步语义。

我的整体评价

APPROVE:上一轮的唯一阻塞项已被实现为显式两阶段交接,并用真实 store 与真实入口的故障注入回归证明可恢复、不重复启动、不额外 spend,同时保留已启动状态的 fail-closed。范围仍与已复现问题相称,未引入新配置或 provider 晋升。剩余 App timer 路径、非 Turn launcher 与打包界面仍按现有 RFC 条件推进;本批准不等于控制面自合并许可之外的额外授权。

English verdict: APPROVE - exact head e5f70a230ba5c75c484f11f168bf88e4fcb33d6a makes a managed start two-phase in the same cadence store, resumes a reserved-but-unstarted start for the same Turn identity once the owner floor is reached, keeps a confirmed start fail-closed, and covers both crash windows with real-store fault-injection regressions; 217 Python cases, 7 TypeScript cadence cases, typecheck, Ruff, mypy, docs governance and the census passed locally.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

上一轮评审在 998c79535711361594a8f141e4809d3111b8f63b 上请求修改,理由是:TS 已原子写入启动预留后,进程可能在 Turn journal 记录 host 尝试之前退出(或 effect 响应丢失),于是同一个 Turn 重启后仍以同一 request ID 申请准入,被 duplicate_or_stale_trigger 永久拒绝——即使等满间隔或给出手动理由也无法恢复,executor 每次都只返回 interval_wait。

本 head 67fd3f9635125bdb9d437759b162c200d3b3368e 修复了该窗口,并合入最新 main。已充分验证的修订是 e5f70a230ba5c75c484f11f168bf88e4fcb33d6a,本 head 相对它只增加一次基础分支合并(effect_runtime_handlers.ts 的 import 冲突按双方新增并存解决),PR 自身产品文件逐字未变,因此复用该修订的验证结论并在本 head 重跑 typecheck 与 cadence 用例。修复方式是让预留与 Turn journal 的交接具备显式、可恢复、可审计的状态,而不是依赖"单层原子写入即安全"的假设。

改动思路

规则所有权不变:TypeScript cadence store 仍是策略与启动记录的唯一 owner,准入在同一 policy 文件锁内完成;Python Turn executor 在任何 host 或 journal 尝试变更之前请求准入,被拒绝时只返回 interval_wait(含下次可执行时间),不调用 host、不写回、不花 quota。

新增的关键区分是"已预留但 host 尚未被持久记录"与"host 尝试已落盘"。启动记录现在带阶段字段:准入写入 reserved,Turn executor 只有在把 host 尝试写入 journal(host_attempt_count 在调用 host 之前先落盘)之后才调用新的 confirm_start 把它置为 started。因此:

  • 同一 request 身份遇到 reserved 记录时,在满足 owner 下限后按原 Turn 恢复(间隔锚点保持首次预留时间,不会被推后或前移);
  • 遇到 started 记录时保持 fail-closed,重复身份与手动理由都不能绕过;
  • 缺少阶段字段的旧记录按"已尝试启动"读取,旧版或手工改写的 store 因此 fail-closed 而不是被当作可恢复预留。

未配置策略时两阶段都不创建 store 或锁文件,默认关闭路径与之前一致。

具体改动

12 文件 +787/−51。TypeScript:automation_cadence.ts 的启动记录增加 reserved/started 阶段并新增 confirmAutomationStart;effect_runtime_handlers.ts 注册 quota.automation_cadence.confirm_start。Python:turn.py 把原先的内联准入闭包提取为可复用的 managed_cadence_start 工厂(准入 + 确认),executor.py 在 host 尝试落盘后、调用 host 前执行确认。测试:TS cadence 用例覆盖恢复/确认/升级/fail-closed;tests/test_loopx_turn_executor.py 新增两个真实 TS store + 真实 Turn 入口的故障注入回归。文档:两份 RFC 记录两阶段交接与恢复语义。本 head 同时合入最新 main(保留双方新增的 effect import/路由,采用 main 已修好的镜像声明,并保留本分支更新的实现基线)。

关键代码讲解

  • admitAutomationStart(loopx/control_plane/quota/automation_cadence.ts):在同一 policy 锁内判定重复/过期触发、owner 下限与手动绕过;同身份的 reserved 记录在满足下限后返回 resumed_unstarted_reservation,started 记录则维持 duplicate_or_stale_trigger。
  • confirmAutomationStart(同文件):幂等确认;无对应预留时返回 reservation_missing 且不创建文件;state 字段缺失的记录在读入时按 started 处理。
  • ManagedCadenceStart / managed_cadence_start(loopx/cli_commands/turn.py):生产 CLI 与回归测试共用同一准入/确认实现,避免测试自建一套旁路语义。
  • _host_result_stage 的确认点(loopx/control_plane/turn_driver/executor.py):host 尝试计数落盘后、host 调用前执行 confirm_start;确认失败即停止,不会在状态未明时启动 host。
  • 故障注入回归(tests/test_loopx_turn_executor.py):分别在"预留后、任何 journal 写入前"与"admission 记录已写、host 尝试计数写入前"中断进程,重启并跨过下限后断言恢复提交、host 仅一次、spend 仅一次、记录最终为 started。

对主干的风险

上一轮点名的 P1 场景现在有真实 TS store + 真实 Python 入口的负例覆盖,两个中断点都验证为可恢复且不重复启动 host、不额外 spend。恢复路径的安全前提是同一 Turn lane 由 journal 锁串行执行、且尝试计数先于 host 调用落盘;这条前提已作为后续风险记录在案,删除该锁或把 host 调用提前都会重新打开重复启动窗口。

仍保持 fail-closed 的部分:已确认启动对同一身份拒绝重复;缺阶段字段的 store 视为已启动;确认失败会以显式失败结束而不是静默继续。范围边界未扩大:App 定时器到 hook、非 Turn launcher、打包设置界面与真实宿主推广仍未验收,本 PR 不晋升任何 provider 或自动化。

验证矩阵(本 head):Python Turn executor/driver/automation cadence/turn envelope/writer fence 共 217 用例通过(含两个崩溃恢复回归);TS cadence 用例 7 项通过;npm run typecheck:control-plane 通过;TS 全量 2936/2961 通过,唯一失败 tests/control_plane_ts/sqlite_capacity.test.ts 在未包含本 PR 的 origin/main 上同样失败(既有环境问题);census 与 goal-instance inventory 9 项通过;docs-governance-smoke.py、配置范围内 Ruff 与 python -m mypy(23 模块)通过。按本 Goal 的 wait_for_ci=false 与维护者要求,本轮未抓取或等待远端 CI;合并前仍对未变 head 运行 merge-readiness。

语义与 CI 对齐

cadence store 由 v1 原地升级为 v2 并新增阶段字段,路径不变,旧版程序拒绝 v2 而不是静默丢弃启动记录;新增的 quota.automation_cadence.confirm_start 属内部 coordination 契约,不改公共 CLI 参数与默认行为。文档已在两份 RFC 中同步语义。

我的整体评价

APPROVE:上一轮的唯一阻塞项已被实现为显式两阶段交接,并用真实 store 与真实入口的故障注入回归证明可恢复、不重复启动、不额外 spend,同时保留已启动状态的 fail-closed。范围仍与已复现问题相称,未引入新配置或 provider 晋升。剩余 App timer 路径、非 Turn launcher 与打包界面仍按现有 RFC 条件推进;本批准不等于控制面自合并许可之外的额外授权。

English verdict: APPROVE - exact head 67fd3f9635125bdb9d437759b162c200d3b3368e makes a managed start two-phase in the same cadence store, resumes a reserved-but-unstarted start for the same Turn identity once the owner floor is reached, keeps a confirmed start fail-closed, and covers both crash windows with real-store fault-injection regressions; 217 Python cases, 7 TypeScript cadence cases, typecheck, Ruff, mypy, docs governance and the census passed locally, and this head re-ran typecheck plus the cadence suite after its base merge.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

上一轮评审在 998c79535711361594a8f141e4809d3111b8f63b 上请求修改,理由是:TS 已原子写入启动预留后,进程可能在 Turn journal 记录 host 尝试之前退出(或 effect 响应丢失),于是同一个 Turn 重启后仍以同一 request ID 申请准入,被 duplicate_or_stale_trigger 永久拒绝——即使等满间隔或给出手动理由也无法恢复,executor 每次都只返回 interval_wait。

本 head a026386e5034301584d27cd4407a5ca5345ed07d 修复了该窗口,并合入最新 main。已充分验证的修订是 e5f70a230ba5c75c484f11f168bf88e4fcb33d6a,本 head 相对它只增加基础分支合并(effect_runtime_handlers.ts 的 import 冲突按双方新增并存解决),PR 自身产品文件逐字未变,因此复用该修订的验证结论并在本 head 重跑 typecheck 与 cadence 用例。修复方式是让预留与 Turn journal 的交接具备显式、可恢复、可审计的状态,而不是依赖"单层原子写入即安全"的假设。

改动思路

规则所有权不变:TypeScript cadence store 仍是策略与启动记录的唯一 owner,准入在同一 policy 文件锁内完成;Python Turn executor 在任何 host 或 journal 尝试变更之前请求准入,被拒绝时只返回 interval_wait(含下次可执行时间),不调用 host、不写回、不花 quota。

新增的关键区分是"已预留但 host 尚未被持久记录"与"host 尝试已落盘"。启动记录现在带阶段字段:准入写入 reserved,Turn executor 只有在把 host 尝试写入 journal(host_attempt_count 在调用 host 之前先落盘)之后才调用新的 confirm_start 把它置为 started。因此:

  • 同一 request 身份遇到 reserved 记录时,在满足 owner 下限后按原 Turn 恢复(间隔锚点保持首次预留时间,不会被推后或前移);
  • 遇到 started 记录时保持 fail-closed,重复身份与手动理由都不能绕过;
  • 缺少阶段字段的旧记录按"已尝试启动"读取,旧版或手工改写的 store 因此 fail-closed 而不是被当作可恢复预留。

未配置策略时两阶段都不创建 store 或锁文件,默认关闭路径与之前一致。

具体改动

12 文件 +787/−51。TypeScript:automation_cadence.ts 的启动记录增加 reserved/started 阶段并新增 confirmAutomationStart;effect_runtime_handlers.ts 注册 quota.automation_cadence.confirm_start。Python:turn.py 把原先的内联准入闭包提取为可复用的 managed_cadence_start 工厂(准入 + 确认),executor.py 在 host 尝试落盘后、调用 host 前执行确认。测试:TS cadence 用例覆盖恢复/确认/升级/fail-closed;tests/test_loopx_turn_executor.py 新增两个真实 TS store + 真实 Turn 入口的故障注入回归。文档:两份 RFC 记录两阶段交接与恢复语义。本 head 同时合入最新 main(保留双方新增的 effect import/路由,采用 main 已修好的镜像声明,并保留本分支更新的实现基线)。

关键代码讲解

  • admitAutomationStart(loopx/control_plane/quota/automation_cadence.ts):在同一 policy 锁内判定重复/过期触发、owner 下限与手动绕过;同身份的 reserved 记录在满足下限后返回 resumed_unstarted_reservation,started 记录则维持 duplicate_or_stale_trigger。
  • confirmAutomationStart(同文件):幂等确认;无对应预留时返回 reservation_missing 且不创建文件;state 字段缺失的记录在读入时按 started 处理。
  • ManagedCadenceStart / managed_cadence_start(loopx/cli_commands/turn.py):生产 CLI 与回归测试共用同一准入/确认实现,避免测试自建一套旁路语义。
  • _host_result_stage 的确认点(loopx/control_plane/turn_driver/executor.py):host 尝试计数落盘后、host 调用前执行 confirm_start;确认失败即停止,不会在状态未明时启动 host。
  • 故障注入回归(tests/test_loopx_turn_executor.py):分别在"预留后、任何 journal 写入前"与"admission 记录已写、host 尝试计数写入前"中断进程,重启并跨过下限后断言恢复提交、host 仅一次、spend 仅一次、记录最终为 started。

对主干的风险

上一轮点名的 P1 场景现在有真实 TS store + 真实 Python 入口的负例覆盖,两个中断点都验证为可恢复且不重复启动 host、不额外 spend。恢复路径的安全前提是同一 Turn lane 由 journal 锁串行执行、且尝试计数先于 host 调用落盘;这条前提已作为后续风险记录在案,删除该锁或把 host 调用提前都会重新打开重复启动窗口。

仍保持 fail-closed 的部分:已确认启动对同一身份拒绝重复;缺阶段字段的 store 视为已启动;确认失败会以显式失败结束而不是静默继续。范围边界未扩大:App 定时器到 hook、非 Turn launcher、打包设置界面与真实宿主推广仍未验收,本 PR 不晋升任何 provider 或自动化。

验证矩阵(本 head):Python Turn executor/driver/automation cadence/turn envelope/writer fence 共 217 用例通过(含两个崩溃恢复回归);TS cadence 用例 7 项通过;npm run typecheck:control-plane 通过;TS 全量 2936/2961 通过,唯一失败 tests/control_plane_ts/sqlite_capacity.test.ts 在未包含本 PR 的 origin/main 上同样失败(既有环境问题);census 与 goal-instance inventory 9 项通过;docs-governance-smoke.py、配置范围内 Ruff 与 python -m mypy(23 模块)通过。按本 Goal 的 wait_for_ci=false 与维护者要求,本轮未抓取或等待远端 CI;合并前仍对未变 head 运行 merge-readiness。

语义与 CI 对齐

cadence store 由 v1 原地升级为 v2 并新增阶段字段,路径不变,旧版程序拒绝 v2 而不是静默丢弃启动记录;新增的 quota.automation_cadence.confirm_start 属内部 coordination 契约,不改公共 CLI 参数与默认行为。文档已在两份 RFC 中同步语义。

我的整体评价

APPROVE:上一轮的唯一阻塞项已被实现为显式两阶段交接,并用真实 store 与真实入口的故障注入回归证明可恢复、不重复启动、不额外 spend,同时保留已启动状态的 fail-closed。范围仍与已复现问题相称,未引入新配置或 provider 晋升。剩余 App timer 路径、非 Turn launcher 与打包界面仍按现有 RFC 条件推进;本批准不等于控制面自合并许可之外的额外授权。

English verdict: APPROVE - exact head a026386e5034301584d27cd4407a5ca5345ed07d makes a managed start two-phase in the same cadence store, resumes a reserved-but-unstarted start for the same Turn identity once the owner floor is reached, keeps a confirmed start fail-closed, and covers both crash windows with real-store fault-injection regressions; 217 Python cases, 7 TypeScript cadence cases, typecheck, Ruff, mypy, docs governance and the census passed locally, and this head re-ran typecheck plus the cadence suite after its base merge.

@huangruiteng
huangruiteng merged commit 5446dea into main Sep 24, 2026
4 checks passed
@huangruiteng
huangruiteng deleted the codex/turn-atomic-cadence-admission branch September 24, 2026 04:56
@huangruiteng

Copy link
Copy Markdown
Collaborator Author

Self-repair and merged decision record (admin-bypass merge, maintainer-authorized).

  • Requested change addressed: the prior review on 998c7953 asked for recoverable admission-to-journal handoff semantics and a real-store fault-injection regression. Both are in this head.
  • Exact head reviewed and merged: a026386e5034301584d27cd4407a5ca5345ed07d; merge commit on main: 5446dea1c15499ec65f0f07118c1b4901be4b9aa.
  • Before the merge, loopx pr-review --check-merge-readiness 4929@a026386e5... returned ready=true with no blocking reasons; git diff a026386e5... 5446dea1c is empty, so merged main equals the validated head.

Repair content:

  • Managed starts are now two-phase in the same cadence store: admission writes a reserved start, and the Turn executor confirms it only after the host attempt is durable in the Turn journal (host_attempt_count lands before the host command runs). The same request identity resumes a reserved start once the owner floor is reached, so a crash between the two no longer strands a Turn; a confirmed start stays fail-closed for the same identity even with a manual reason, and a store record without the phase field is read as an attempted start.
  • The CLI admission closure became a module-level managed_cadence_start factory shared by production and tests, so the regression exercises the same wiring instead of a test-only path.
  • Two new regressions interrupt a real TypeScript-store-managed start: (1) after the reservation and before any journal write, (2) after the admission record but before the attempt counter. Both restart, cross the floor, commit, start the host exactly once and spend exactly once, ending with a started record.
  • Latest main was merged; conflicts were import-block only (keep both sides' effect imports and routes) plus the RFC mirror declaration, where main's repaired wording is used while keeping this branch's implementation baseline.

Validation at the merged head:

  • Passed: 217 Python cases across the Turn executor (including both crash-injection recoveries), Turn driver, automation cadence, turn envelope and legacy writer fence modules; 7 TypeScript cadence cases; npm run typecheck:control-plane; configured Ruff; mypy (23 modules); examples/docs-governance-smoke.py; census and goal-instance inventory (9 cases); change-quality receipt cqr_683f0665229977ecddbc (decision=pass).
  • TypeScript full suite: 2936/2961 on the reviewed revision whose product files are byte-identical to this head. The single failure, tests/control_plane_ts/sqlite_capacity.test.ts ("small capacity entrypoint exercises real SQLite..."), reproduces identically on unmodified origin/main.
  • Not awaited: remote CI, per this Goal's wait_for_ci=false and the maintainer's explicit instruction for this run.
  • Carried limits: the App timer-to-hook path, non-Turn launchers, packaged settings UI and real host promotion remain unqualified; resume safety depends on the Turn journal lock serializing one lane and on the attempt counter landing before the host runs, which is recorded as a standing risk in the receipt.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant